Skip to main content

Over recent months, a growing share of our incident response engagements has had one thing in common: somewhere in the attack chain, the threat actor had been talking to a large language model (LLM). The calls themselves still arrive the way they always have, often late on a Friday, from someone who is calm in the way people are when they are very much not calm, and whose first sentence is some variation of “we think something’s happened”. What has changed is what we find once we start looking.

That isn’t a prediction. It’s what we’ve found in the artefacts. Below are five patterns we’ve seen repeatedly across our engagements, and what they mean for defenders. 

James Thoburn
James Thoburn

Director, Incident Response 

jthoburn@thomasmurray.com

1. The intruder left the chat log behind

The most striking evidence is also the most mundane. On several engagements this year we have recovered, from staging directories and browser profiles on compromised hosts, the conversation history between the threat actor and an LLM as the intrusion progressed. Not a fragment. The whole thing.

The transcripts read like a junior operator being coached through an engagement by a very patient senior. “This command failed with the following error, what should I try?” “How do I find the backup server on this network?” “Write me a PowerShell one-liner that does X without triggering Y.” The model obliges, the operator pastes, the operator reports back. Rinse and repeat until they have domain admin.

Two things follow from this. First, the barrier to entry has dropped considerably. Individuals who would not previously have been able to work through a Windows domain unaided now can, with a tireless assistant that never sleeps and never gets bored. Second, and more usefully for us, these transcripts are invaluable for attribution and scoping. They tell us what the actor knew, what they wanted, where they got stuck and, crucially, what they tried that didn’t work. If your incident response provider isn’t looking for them, it should be.

2. AI-written malware, and why you shouldn’t pay for a broken product

It sounds strange to say, but there was a time when ransomware was reliable. Groups such as LockBit ran a mature operation: the encryptor worked, the decryptor worked, and the negotiation was, in its grim way, a business transaction with a predictable outcome. You paid, you got your data back, and the whole thing was engineered to make paying feel like the sensible choice.

That reliability is eroding, and we think AI-assisted development is a significant part of the reason.

What we are now seeing in the wild is encryption tooling that has clearly been assembled quickly, by someone who doesn’t understand the environment it is being deployed into. Encryptors that don’t stop running processes before they start, so database files are left half-encrypted and unrecoverable even with the key. Encryptors that walk straight into active backup jobs and corrupt the backup mid-write. Encryptors that have no concept of a virtualised estate and encrypt VMDKs, datastores and snapshots indiscriminately, taking down 30 virtual machines to hit one target. And decryptors that simply don’t work, because nobody tested the round trip.

The result is that the damage is frequently far worse than it needed to be for the actor’s purposes and, more importantly, the ransom buys you nothing. We have now seen multiple cases where a victim paid and received a decryptor that either failed outright or recovered only a fraction of the estate. When the criminal’s product is this unreliable, the traditional argument for payment collapses. It was never a good argument, but it used to at least be a coherent one.

Our advice has shifted accordingly: before any conversation about payment, get a technical assessment of whether the encryptor actually did what it claimed. Increasingly, the answer is no.

3. Cloud workers and the ten-minute phishing site

Business email compromise remains the most common reason we are called out, and the front end of it has become alarmingly slick.

Serverless platforms such as Cloudflare Workers were built to let developers deploy code to the edge in seconds, on trusted domains, with a free tier and no infrastructure to manage. That is exactly what a phishing operator wants too. We are seeing credential harvesting pages deployed on these platforms at scale: pixel-accurate clones of Microsoft 365 and Okta login flows, complete with correct branding, multi-factor authentication (MFA) prompts, error handling and a redirect to the real site afterwards, so the victim never notices.

This is where AI has changed the economics. Building a convincing credential page used to require some skill and some time. Now the operator describes what they want, gets working code back, deploys it to a worker and has a live lure on a reputable-looking domain in the time it takes to make a cup of tea. When one is taken down, the next is already up. Domain reputation, the tool most email filters lean on hardest, is close to useless when the domain belongs to a legitimate content delivery network.

For defenders, the practical implications are clear. Phishing-resistant MFA (FIDO2 or passkeys) is no longer a nice-to-have. Conditional access policies need to be tight enough that a stolen session token is worth less than it currently is. And your users need to understand that “the link went to a real-looking page” is exactly the point.

4. Custom exfiltration tooling, and the limits of hunting for rclone

For a long time, data theft ahead of extortion followed a recognisable pattern. The actor drops rclone or a similar utility, points it at a MEGA or other cloud storage account, and off it goes. Defenders learned to watch for those binaries, those hashes and those command lines, and it worked, up to a point.

That point has arrived. On recent engagements we have recovered bespoke exfiltration tools that appear to have been written for the job: small, compiled, free of external dependencies and built on well-established protocols such as S3, so the traffic looks like ordinary cloud backup activity, because in every technical sense it is. They chunk, they compress, they throttle to stay under egress alerting thresholds, and they clean up after themselves. Several bore the hallmarks of AI-assisted development: sensible structure, verbose comments and the kind of tidy error handling that a hurried human writing a throwaway tool would never bother with.

None of this is exotic. It is simply that a custom tool has no signature, and a fresh compile has no hash reputation. The lesson is one the industry has been slow to learn: hunt for the behaviour, not the binary. Large outbound transfers to newly seen cloud storage endpoints, unusual volumes from file servers outside business hours, processes reading across many shares in a short window. These are detectable regardless of what wrote the code, and they are the signals that matter.

5. The IT helpdesk that calls you

The final pattern is the least technical and, in our experience this year, the most effective.

Where an organisation’s Microsoft Teams tenant allows communication with external tenants (and a surprising number do, often by default and without anyone having made a deliberate decision), threat actors are simply calling staff directly. The caller presents as IT support, often with a plausible display name and some context lifted from LinkedIn or a previous breach. They say there is an issue with the user’s account, or a mandatory security update, and ask the user to install a remote support tool, approve an MFA prompt or read out a code.

It works because it exploits the one thing every security awareness programme has spent years reinforcing: when IT asks you to do something, you do it. The call is often preceded by an email bombing campaign, so the “IT” call arrives just as the user is drowning in spam and desperate for help. With AI in the loop, the pretext is well written, the follow-up emails are polished and the whole operation scales in a way that a human social engineer working a phone never could.

The fix here is largely a configuration change. External access in Teams should be restricted to named partner domains, or disabled outright unless there is a documented reason for it. Staff need to know that IT will never cold-call them and ask them to install anything, and they need a way to verify a caller that doesn’t rely on the caller. It is not a complex control, which is exactly why it is so frustrating to see it missing after the fact.

What this adds up to

None of the five patterns above is new in kind. Intrusions, ransomware, phishing, exfiltration and social engineering have been the staples of incident response for as long as the discipline has existed. What has changed is speed and reach. The tooling is built faster and deployed faster, by people with less skill, in greater volume and with less care for the collateral damage they cause.

That last point deserves emphasis: the threat actors who have leaned on AI in recent months were not more sophisticated. In most cases they were less so. But they were more numerous, they moved more quickly, and when their tools broke, they broke in ways that hurt the victim rather than the attacker.

For defenders, the response is mostly the unglamorous work: behavioural detection over signatures, phishing-resistant authentication, locked-down collaboration tenancies, tested and isolated backups, and an incident response capability that knows what AI-assisted intrusions look like in the artefacts. None of it will stop the Friday afternoon phone calls. But it might make the next one a lot shorter.

Cyber Risk

Incident Response

Thomas Murray’s incident response team is trained to respond quickly and efficiently to incidents and help your business get back on track.

Learn more